Trends in Hearing
○ SAGE Publications
Preprints posted in the last 90 days, ranked by how well they match Trends in Hearing's content profile, based on 15 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Benecke, J.; Whitmer, W. M.
Show abstract
In conventional hearing-aid personalisation, clinicians cannot hear what their patients hear, and patients cannot often reliably detect or describe what they hear. Self-adjustment avoids this issue but requires user controls that adjust hearing-aid signal processing parameters to be effective, efficient and easy. In this study, we explored (a) the roles of interface complexity and stimulus type in the self-adjustment of hearing-aid gain, and (b) how well individuals can adjust one sound to match another to assess the same interfaces and stimuli. Adult hearing-aid users with mild to moderate symmetrical sensorineural hearing loss repeatedly adjusted the gain (a) to their preference from individual prescription (n = 41) and (b) to match their previous preferences from a random starting point (n = 32) using three interfaces representing different bass/mid/treble configurations and three stimuli (music, speech and speech-in-noise). The large interindividual variability in self-adjusted gains clustered into three patterns of deviation from initial prescription: increased relative bass, overall gain reduction, and close to initial prescription. There were no substantial effects of interface nor stimulus on self-adjustment reliability (median {sigma} = 2.8 dB), whereas absolute sound-matching error increased with increasing interface complexity and centre frequency. Neither individual matching accuracy nor questionnaire responses predicted either self-adjusted gains or reliability. Overall, these results show that many - but not all - hearing-aid users can adjust gains with reasonable reliability, and while it can be difficult to predict the behaviour from the individual, the individual applies a similar self-adjustment behaviour across different interfaces and stimuli.
Colak, H.; Guo, X.; Benzaquen, E.; Gurusiddappa, M.; Banerjee, A.; Choi, I.; Sedley, W.; Griffiths, T. D.
Show abstract
ObjectivesOutcomes following cochlear implantation vary substantially across adult recipients, and the cognitive and perceptual factors contributing to this variability are not fully understood. This poses a challenge for developing strategies to improve cochlear implant outcomes, as such approaches require a clearer understanding of the mechanisms underlying individual listening difficulties. In this study, we investigated auditory cognitive measures in cochlear implant (CI) users to further elucidate the origins of this variability. DesignThirty-seven adult cochlear implant users completed measures of auditory cognition, comprising auditory working memory (AWM) and sound segregation ability, measured using an auditory figure-ground task (AFG), as well as measures of peripheral temporal and spectral processing, comprising the temporal modulation detection threshold (TMDT) and spectral ripple discrimination threshold (SRDT). Speech perception outcomes were assessed using word-in-noise (WIN) and sentence-in-noise (SIN) tasks. Separate multiple linear regression models evaluated the unique contribution of the auditory cognition measures to WIN and SIN performance, after accounting for the peripheral measures. ResultsBoth regression models explained a substantial proportion of variance in speech-in-noise outcomes (WIN: adjusted R{superscript 2} = 0.55; SIN: adjusted R{superscript 2}=0.57, both p < 0.001). For WIN performance, AFG and AWM were significant predictors. A similar pattern was found for SIN performance, where lower AWM ability and poorer AFG segregation were linked to poorer sentence listening in noise. No significant effects of spectral ripple discrimination or temporal modulation detection were observed in either model, even though both were significantly correlated with WIN performance. ConclusionsThese findings indicate that auditory working memory and sound segregation ability are robust predictors of speech-in-noise outcomes in adult cochlear implant users, across both word- and sentence-level measures. Together, the results may help explain why speech-in-noise outcomes remain highly variable among CI users, even when basic sensory encoding abilities are taken into account. Incorporating measures of auditory working memory and fundamental sound segregation may therefore improve outcome prediction and help in developing more individualised rehabilitation strategies.
Fish, E.; DiNino, M.
Show abstract
Acoustic cues such as pitch and spatial location allow listeners to attend to a target speaker and ignore competing talkers, aiding speech recognition in background noise. Diminished ability to utilize acoustic cues for speech stream segregation may thus contribute to older adults' challenges hearing in noise. Adults aged 18-74 completed a speech-in-speech identification task with three conditions containing 1) only pitch cues (fundamental frequency), 2) only spatial cues (interaural time differences; ITDs), and 3) both pitch and spatial cues for segregating a target talker from competing talkers. Hearing thresholds at standard and extended high frequencies (EHFs), auditory brainstem responses (ABRs), and digit span scores were acquired to examine the influence of sensory and cognitive factors on use of each acoustic cue for speech-in-speech recognition. Significant differences were observed between cue condition scores indicating that use of the available cue(s) drove performance. ABR metrics were not a significant predictor but digit span scores significantly predicted scores on all three cue conditions. Working memory abilities therefore set a baseline for participants' speech-in-speech recognition regardless of the acoustic content. Hearing thresholds at standard frequencies significantly predicted scores on the Pitch condition. EHF hearing thresholds better predicted Spatial and Both Cue condition performance, suggesting that EHF thresholds represent auditory processing important for coding ITDs. Age group analysis revealed that older adults (aged 40+) performed significantly more poorly on all cue conditions of the speech-in-speech recognition task relative to younger adults. Age-related changes in auditory sensory processing may therefore impair older adults' speech-in-noise perception by reducing their ability to use acoustic cues for segregating target and competing speech.
Delaram, V.; Ananthanarayana, R. M.; Trine, A.; Miller, M. K.; Stecker, G. C.; Buss, E.; Monson, B. B.
Show abstract
Several types of cues contribute to speech recognition in multi-talker environments. In this study, we investigated how talker head-orientation related (THOR) cues and extended high- frequency (EHF; >8kHz) cues affect speech-in-speech recognition for both female and male speech. We examined the THOR benefit associated with a non-facing masker talker head orientation (relative to a facing orientation) as a function of masker talker facing angle. The target talker always faced the listener, whereas co-located maskers were tested with eight different masker head angles, ranging from 0{degrees} (facing the listener) to facing 180{degrees} away. Two filtering conditions were tested: full- band and low-pass filtered at 8 kHz. A THOR benefit was observed at masker head angles greater than 45{degrees}, increasing from 2 dB to 8 dB between angles of 67.5{degrees} and 180{degrees}. This benefit was reduced for low-pass filtered speech. Access to EHF cues improved performance, but only for masker head angles >22.5{degrees}. There was no significant relationship between 16-kHz pure-tone thresholds and performance for young, normal-hearing listeners with good EHF hearing. These findings indicate that listeners benefit from non-facing masker talker head orientations >45{degrees} when the target talker is facing the listener, with greater benefit for larger head angles.
Richardson, B. N.; Guru Adimurthy, M.; Brown, C. A.; Ihlefeld, A.; Rosen, M. J.; Shinn-Cunningham, B. G.
Show abstract
Intelligible speech disrupts selective auditory attention more than an unintelligible stream. However, low-level acoustic features of intelligible speech are relatively similar to target speech, confounding results. While controlling acoustic similarity and limiting energetic masking, we examined how masker intelligibility affects behavior and electroencephalography (EEG). Normal hearing listeners detected color words within a target stream of randomly timed words while ignoring an ongoing masker. Maskers were either spoken by the same or a different talker and comprised either isochronous sequences of intelligible words or temporally scrambled versions. Scrambled maskers either lacked broadband energy changes over time (Experiment 1) or were amplitude modulated to have the same energy profiles as intelligible, isochronous maskers (Experiment 2). In both experiments, scrambled maskers yielded better performance than intelligible maskers. For intelligible maskers, performance was better for different compared to identical talkers. EEG responses paralleled behavior: target-evoked onset responses were larger for scrambled than for intelligible maskers, particularly for identical talkers. Later target recognition responses were larger for color than other target words but unaffected by masker type or talker. Even when low-level acoustic features were carefully matched, intelligible maskers impaired auditory attention and reduced target-evoked neural responses more than scrambled maskers, implicating early sensory filtering.
Benecke, J.; Whitmer, W. M.
Show abstract
Studies employing and evaluating hearing-aid self-adjustments frequently label the resulting settings as 'preferences' without questioning that label. There is a lack of inquiries into the rationale behind participants' choices or their approach to self-adjustment. We here employed a think-aloud protocol concurrent with individuals' self-adjustments to capture qualitative perspectives on these processes. Thirty-six adult hearing-aid users were asked to verbalise what they were thinking while adjusting to their preference three stimuli (music, speech or speech in noise) using three control interfaces (1 slider jointly controlling bass/treble, 2 sliders separately controlling bass and treble or 3 sliders controlling bass, mid and treble). Audio-video recordings were transcribed, annotated and analysed using inductive content coding and thematic analysis. Participants' approach and navigation of the self-adjustment process was best defined as a continuum between exploratory and anticipatory archetypes, with both types sometimes occurring at different stages within the same adjustment. The exploratory type initially assesses control functions, then selects the optimal setting from available options, and evaluates sound changes affectively in isolation or comparatively against prior settings. The anticipatory type begins by evaluating the sound, identifying issues, adjusting settings guided by past experiences, and evaluating adjustments against the identified issues. Additionally, the descriptors elicited during adjustments generally agreed with expert-based data. The resulting qualitative framework of self-adjustment helps explain both how hearing-aid users navigate the personalisation process and how self-adjustment can promote greater understanding of the options available and greater ownership in that process.
Dirks, C. E.; Guest, D. R.; Oxenham, A.
Show abstract
Context effects are ubiquitous across sensory systems and reflect a general encoding principle for both simple and complex stimuli. One simple context effect, contraction bias, manifests in two-interval perception tasks as a bias of the perceived magnitude of the first stimulus toward the center of the overall magnitude range. The underlying cause of contraction bias is unclear. One explanation is that a listeners magnitude estimate of the first stimulus is combined with a perceptual anchor, usually the mean stimulus magnitude, biasing it toward the anchor (sensory model). An alternative explanation is that a listeners response criterion shifts, based on the magnitude of the stimulus pair, relative to the mean magnitude of the stimuli range (decision model). Two pitch-discrimination experiments were performed to test these hypotheses in the auditory domain. The first was a forced-choice discrimination task, where listeners were asked to identify the higher or lower tone in a pair. The second was a same-different task where listeners indicated whether or not the two tones in a pair differed in frequency. Contraction bias was observed in the higher-lower discrimination task, even after extensive perceptual training with feedback. In contrast, no contraction bias was observed in the same-different task. Computational models of the sensory and decision hypotheses were fit to data from both experiments. The sensory model captured the pattern of results the higher-lower experiment but erroneously predicted a contraction bias in the same-different task. The decision model produced similar predictions to the sensory model in the higher-lower task but correctly predicted no contraction bias in the same-different task, and produced lower prediction errors and more stable parameter estimates in both paradigms. Overall, the results suggest that the underlying nature of the contraction bias may reflect decision, rather than sensory, biases based on the context.
Dantanarayana, N. D.; Li, Y.; Litovsky, R. Y.; Borjigin, A.
Show abstract
Humans often communicate and learn in noisy, complex listening environments. Here, we investigated the effects of spatial hearing and semantic context cues on speech intelligibility and listening effort in young adults with typical hearing. The listening task included conditions in which target speech and speech maskers were either spatially co-located or separated. Target sentences were either semantically coherent or anomalous, while the masker comprised a mixture of two coherent sentences. Results showed higher speech intelligibility in spatially separated than co-located conditions, demonstrating a robust spatial release from masking (SRM), which is consistent with prior findings. SRM did not differ between semantically coherent and anomalous sentences, indicating comparable benefits of spatial cues across semantic contexts. However, within each spatial configuration, intelligibility was higher for coherent than anomalous sentences. Listening effort, indexed by peak pupil dilation in pupillometry measurement, was reduced in spatially separated conditions, suggesting a trend toward a release from listening effort. Analysis of the timing of peak pupil dilation revealed a significantly delayed peak dilation for anomalous sentences in the co-located condition compared with coherent sentences in the separated condition, indicating increased processing demands in the absence of spatial and semantic cues. Finally, SRM was correlated with the magnitude of release from listening effort for coherent sentences, but not for anomalous sentences, suggesting that intelligibility and listening effort benefits might co-occur when contextual cues are available.
Chao, M.; Holloway, C. A.; Miller, L. M.; Mankel, K.
Show abstract
Difficulties understanding speech in noise remain a common complaint even among listeners with normal hearing sensitivity, highlighting the need for objective, more effective measures of real-world listening. The goal of this study was to validate the use of a novel, chirped-speech (Cheech) stimulus - continuous, naturally-spoken speech fused with chirps designed to elicit robust auditory evoked potentials - to characterize relationships between speech recognition, listening effort, and auditory neural encoding. Twenty-five normal-hearing adults completed a sentence-recognition task using both original (unmodified) and Cheech-modified AzBio sentence lists in quiet, +3 dB, and -3 dB signal-to-noise ratio (SNR) conditions while neural responses from the brainstem through cortex were recorded simultaneously. Speech recognition remained near ceiling in quiet but declined with decreasing SNR for both original and Cheech stimuli. Compared with clean speech, Cheech-modified speech showed slightly poorer recognition performance as SNR decreased and somewhat higher perceived effort overall. Yet, Cheech was highly effective at evoking auditory responses from the brainstem (auditory brainstem response, ABR) through the cortex (including middle- and late-latency responses, MLR and LLR) even with <5 minutes listening time per condition. Neural responses showed reduced amplitudes and prolonged latencies as SNR decreased. In general, ABR latencies and wave I amplitudes were associated with speech-in-noise recognition performance, whereas cortical responses (MLR Na, Nb, and LLR P1) were associated with subjective workload. These findings show that Cheech-modified speech preserves intelligibility while yielding robust, multilevel neural recordings during sentence perception, offering a promising approach to examine hierarchical auditory processing under ecologically relevant speech-in-noise conditions.
Wang, Z.; Li, G.; Yu, Y.; Wu, J.; Yu, Z.; Meng, Y.; Wang, S.; Dong, C.
Show abstract
Efficient face-to-face communication relies on the integration of auditory speech and visual articulatory signals. Over the past five decades, the McGurk illusion has been widely used as an index of audiovisual speech integration. However, substantial variabilities in susceptibility to the illusion across participants and speakers limit its reliability as a stable measure of audiovisual integration ability. Here, we introduce the McGurk illusion dataset (MID), which, to our knowledge, is the largest publicly available McGurk stimulus dataset to date. The MID comprises auditory (N = 400), visual (N = 400), and audiovisual (N = 640) speech stimuli generated from 80 Mandarin speakers and validated through behavioral judgments across 360,900 trials. Using this dataset, we characterized the acoustic and facial articulatory properties of McGurk stimuli, replicated substantial inter-participant and inter-speaker variabilities in illusion susceptibility, and revealed the associations between variations in McGurk illusion rate and the variations in unisensory perception, audiovisual correspondences, and speakers characteristics. Furthermore, the stimulus set enabled systematic comparisons of the reliability of different McGurk illusion-based indices of audiovisual speech integration. Overall, the MID not only provides a standardized resource for investigating audiovisual speech integration and its alterations across populations, but also supports research on speaker normalization, lip-reading, and speech perception.
Vickers, D.; Buelt, L.; Arram, E.; Picinali, L.; Salorio-Corbetto, M.; Chowdhury, K.; Clarke, C.; Freemantle, N.; Jiang, D.; Parmar, B.; Early, F.; Driver, S.; Bordea, E.; Hill, T.; Cullington, H.; Kukiewicz, F.; Rocca, C.; Kitterick, P.; Corbett, F.; Nightingale, R.; Blackstone, J.; Ahmed, N.; Somerset, S.; Van Zalk, N.; Mahon, M.
Show abstract
Introduction Deafness is the most common human sensory deficit. Cochlear implantation is the primary intervention for severe-to-profound deafness. Currently, over 7000 people have bilateral cochlear implants (CIs) in the United Kingdom (UK), most of whom are children. Patient feedback suggests that for children with bilateral CIs, everyday communication is challenging and tiring, with extra effort required to integrate information from two ears, especially in noise, and that current rehabilitation techniques are not engaging, or appropriate to their lifestyles. To address these issues, researchers developed the Both EARS (BEARS) training package comprised of three virtual reality games to improve sound localization and listening in noise. Objectives This protocol describes the design and methodology of a multi-center phase III randomized controlled trial (RCT) to evaluate whether use of the BEARS training package alongside usual care compared to only receiving usual care improves speech-in-noise perception, hearing experiences, vocabulary and quality of life and reduces listening effort in children and young people (aged 8 -16 years (inclusive) with bilateral CIs. Methods This RCT is currently underway in 16 clinical CI departments in National Health Service or university hospitals across the UK. The intervention involves 3 months of spatial-listening training delivered via the BEARS training package in addition to any routine rehabilitation. The control is usual care (routine rehabilitation clinical care pathway). The primary outcome is the difference between the intervention groups in speech-in-noise perception score at 3 months derived from the spatial speech in noise (SSiN-VA) test. Recruitment closes at the end of the day on 31st July 2026, and end of data collection is 31st October 2026. Data analyses will be reported by 31st March 2026. Significance This is the largest known trial of children and young people with bilateral CIs. It will generate high-quality evidence on speech-in-noise outcomes and inform training interventions to improve spatial listening. Trial registration ClinicalTrials.gov registration: NCT05808543; UKs clinical study registry (ISRCTN92454702)
Plegat, M.; Araujo Vitoria, M.; Marinato, G.; Tita, B.; van der Lans, C.; Pijfers, M.; Esposito, M.; Bertovic, M.-S.; Formisano, E.; Giordano, B. L.
Show abstract
Natural-sound research requires stimulus sets that combine acoustic standardization with detailed behavioural characterization. We present 1,377 two-second sounds representing 240 expert-defined source--action classes. We call this database "MaMa Sounds", as it resulted from the collaborative effort of two academic teams in Maastricht and Marseille. The sounds were manually curated, segmented, sampled at 16 kHz, and labelled with a noun identifying the source and a verb identifying the action. We release deidentified trial-level identification and familiarity data together with multiple per-sound norms (e.g., identification accuracy, confidence and agreement; familiarity), along with overall norms derived with principal component analysis. Noun, verb, and joint noun--verb norms are provided as direct means and medians with the number of contributing observations. This battery preserves process-specific information, while two principal-component scores provide compact overall behavioural-identifiability measures derived from response ease, semantic correspondence, agreement, and familiarity. The repository also contains deterministic response-cleaning code, participant and reference Word2Vec representations, and code reproducing the public sound-level tables. The resource supports stimulus selection, matching, and continuous modelling in auditory cognition and neuroscience.
Conner, A. N.; Mondul, J. A.; Kulkarni, S.; Mackey, C. A.; Batchu, A.; Temghare, N.; Hackett, T. A.; Ramachandran, R.
Show abstract
Noise exposure can produce lasting auditory dysfunction in the absence of permanent threshold shifts or hair cell loss, yet the functional consequences of temporary threshold shift (TTS) remain poorly defined in translational models. We assessed auditory brainstem responses (ABRs) and distortion product otoacoustic emissions (DPOAEs) in rhesus macaques (n = 13) at 2 and 9-10 months following a single moderate noise exposure that induced TTS. Previous histological analyses of these macaques showed no significant loss of hair cells or ribbon synapses but revealed persistent broadening of inner and outer hair cell ribbon-volume distributions. After exposure, DPOAE amplitudes and thresholds and ABR thresholds returned to pre-exposure values and showed low-frequency enhancement at later time points. Suprathreshold click- and tone-evoked ABR amplitudes were largely preserved or enhanced after exposure, consistent with compensatory gain. In contrast, macaque-specific chirp-evoked ABRs showed modest amplitude reductions and latency prolongation across waves, indicating altered neural synchrony at standard stimulus presentation rates, but with variable time courses. More temporally demanding paradigms revealed persistent impairments. ABRs to faster click rates and shorter paired-click intervals showed reduced adaptability in response amplitude and timing after normalization, with deficits persisting through 9-10 months. Increased inner hair cell ribbon-volume variability was more consistently associated with temporal response measures, including latency, paired-click recovery, and rate adaptation, than with amplitude-based ABR measures. Together, these findings reveal a lasting dissociation between response magnitude and fidelity after TTS: suprathreshold responses may be preserved or enhanced, while neural synchrony and temporal adaptability remain impaired. Increased presynaptic ribbon volume variability may serve as a structural marker of synaptic remodeling accompanying hidden auditory dysfunction, rather than as a direct determinant of suprathreshold response magnitude. Temporally demanding ABR paradigms may supplement threshold-based diagnostics for detecting persistent noise-induced auditory dysfunction.
Laird, E. C.; Gosbell, D.; Dall'Est, A.; Malicka, A.
Show abstract
Objective: To evaluate the efficacy, engagement, and usability of Tune Out, an unguided, self-paced online tinnitus management program, for reducing tinnitus severity in adults with tinnitus. Design: A two-arm, parallel-group randomised controlled trial was conducted with Australian adults reporting diagnosed or self-reported tinnitus. Participants were randomised to immediate access to Tune Out or a waitlist control group. Outcomes were assessed at baseline, 6 weeks, and 12 weeks. The primary outcome was tinnitus severity measured using the Tinnitus Functional Index (TFI). Secondary outcomes included tinnitus handicap, psychological symptoms, program engagement, self-efficacy, and usability. Results: Eighty-eight participants were randomised: 43 to the intervention group and 45 to the waitlist control group. The primary outcome analysis included 63 participants at 12 weeks. A significant Group x Time interaction was observed for TFI total score, indicating greater reductions in tinnitus severity over time in the intervention group compared with waitlist control, F(2, 102.57) = 5.95, p = .004, partial 2= .104. Significant effects were also observed for tinnitus handicap, F(2, 106.76) = 4.12, p = .019, partial 2 = .072. Effects on psychological symptoms were less consistent, although anxiety showed a significant Group x Time interaction, F(2, 116.85) = 3.63, p = .030, partial 2 = .059. At 12 weeks, 23.1% of intervention participants achieved a clinically meaningful reduction in tinnitus severity compared with 5.4% of controls. Program use was highly variable, with a median use of 1.10 hours, and 25.6% of intervention participants recording no use. Usability ratings were favourable among respondents, with a mean System Usability Scale score of 73.13. Conclusions: Tune Out demonstrated preliminary efficacy for reducing tinnitus severity and tinnitus handicap compared with waitlist control. Effects on broader psychological symptoms were less consistent. Although usability was rated positively, low and variable engagement highlights the need for strategies to support uptake and sustained use in unguided digital tinnitus interventions.
Maidment, D. W.; Habib, A.; Gomez, R.; Benton, C.; Ferguson, M. A.
Show abstract
The availability of hearing aids that can connect wirelessly to smartphone technologies via Bluetooth has grown exponentially in recent years. However, there is limited evidence assessing the benefits of user-adjustability afforded by these devices. This study aimed to assess the benefits of smartphone-connected hearing aids and an accompanying application (or app) in new and existing hearing aid users. In this single-centre, prospective, observational study, 44 adult hearing aid users (14 new and 30 existing) were recruited. Participants were fitted bilaterally with smartphone-connected hearing aids that could be adjusted by the user via an app. Self-reported outcome measures were collected at fitting and after seven-weeks of using the device in everyday life. For both new and existing hearing aid users, significant improvements in social participation, hearing-related fatigue, and hearing aid benefit and satisfaction were found. For existing hearing aid users, all outcomes were significantly better for the smartphone-connected hearing aids plus app in comparison to their existing hearing aids that did not connect to a smartphone, all with moderate-to-large clinical effect sizes (d> .6). User-controllability via the app was identified as the key benefit, and most participants (68%) reported that the app met their needs 'extremely' or 'very well'. These results suggest that, when used in conjunction with an app, smartphone-connected hearing aids can improve hearing outcomes due to greater user-controllability to improve listening. Thus, smartphone-connected hearing aids have the potential to facilitate patient-centred care, empowering the individual to successfully manage their hearing loss.
Axe, D.; Muthaiah, V. P. K.; Farhadi, A.; Heinz, M. G.
Show abstract
Sensorineural hearing loss can result from different pathologies, but the primary diagnostic method is a threshold-based audiogram, which is insensitive to some forms of cochlear dysfunction. Individuals may experience difficulty understanding speech in noise despite normal audiometric thresholds. Because most cochlear insults damage both inner (IHCs) and outer hair cells (OHCs), the contribution of IHC dysfunction to auditory-nerve coding has been difficult to isolate. We used the IHC-selective ototoxicity of carboplatin in chinchillas to examine how IHC dysfunction, with preserved OHC function, affects temporal-envelope coding in auditory-nerve fibers (ANFs). Carboplatin produced 10 to 20% IHC loss with stereocilia damage in surviving IHCs, while OHC-dependent measures such as DPOAEs and ANF thresholds were unchanged. Suprathreshold ABR wave 1 was reduced, whereas wave 5 was preserved, suggesting central compensation. Both spontaneous and driven firing rates decreased following exposure. Mean vector strength to amplitude-modulated tones was unchanged, but response variability increased. Neurometric analysis and mutual information showed degraded AM detection in carboplatin-exposed fibers, an effect accounted for by reduced driven rate (i.e., normalizing spike counts across groups removed the group difference). Background noise degraded AM coding similarly in both groups. Pooled-neurometric modeling showed that population redundancy compensated for impaired fibers in quiet, but not in noise, where carboplatin-exposed pools remained worse. These findings indicate that IHC dysfunction degrades envelope coding by reducing neural output rather than by altering temporal synchrony. This study suggests IHC dysfunction is a phenotype consistent with "hidden hearing loss" (but distinct from cochlear synaptopathy), and motivates suprathreshold clinical assays.
Marrone, J. P.; Ziliak, M. C.; Bartlett, E. L.
Show abstract
Auditory brainstem responses (ABRs) are a core part of objective functional evaluations of hearing sensitivity and subcortical auditory transmission. Manual assessments of ABR waveforms are still a primary means by which thresholds and peak amplitudes and latencies are measured, which is time-consuming and prone to user variability. Automated methods have offered promising alternatives for ABR classification, but they have sometimes been limited in accuracy or robustness. Here, we developed and tested a supervised convolutional neural network (CNN) based ABR peak classifier that works across sound levels and sound frequencies that can be run quickly on a personal computer using single or dual-channel ABR inputs. For ABR peaks I, III, IV, and V, the classifier achieved over 95% accuracy. High accuracy was maintained even after noise-exposure causing temporary or permanent threshold shifts, and over 90% of peaks were within 0.041 ms (1 sample) of the manually identified peak. Only a few hundred samples were needed to train the network, making it widely amenable to smaller data studies or where the number of subjects or sessions may be low.
Scott, M. T.; Limon, P. N.; Popelka, G. R.; Butts Pauly, K.; Norcia, A. M.; Ash, R. T.
Show abstract
Auditory confounds have proven to be a major hurdle in the elucidation of veridical neuromodulation effects with transcranial ultrasound stimulation (TUS). Auditory noise masks are an essential method to reduce the audibility of TUS and have shown promise in several studies. Here we describe a novel approach for design, calibration, and psychometric validation of auditory noise masks to reduce the perceptibility of TUS. White noise masks and spectrum-tuned masks matched to a TUS protocol that generates highly salient auditory costimulation (487.5 Hz pulse repetition frequency, 10% duty cycle, 68 W/cm2 pulse-peak average intensity, 500 kHz acoustic frequency) were generated, and dB(A) levels were calibrated with an artificial ear. The masker levels needed to reduce TUS detection performance in a two-interval forced choice task were determined with an adaptive QUEST+ staircase in 20 neurotypical participants. Detection performance of this highly salient TUS protocol was driven to near chance performance (<55%) in 17/20 participants with white-noise and 18/20 participants with spectrum-tuned noise. However, high masker levels approaching safety limits were needed to render TUS inaudible for the majority of participants, indicating the need for formal masker calibration for these TUS settings. Additionally, against expectation the spectrum-tuned masker did not significantly outperform the white-noise masker, suggesting that perceptibility of TUS auditory costimulation does not lawfully follow the sound expected from its pulse envelope and the known spectrum of human hearing.
Davies, T.; Bleeck, S.
Show abstract
Objective: This study investigated whether plosive consonants carry a perceptual loudness weighting that significantly exceeds that of non-plosive consonants when judged by hearing-impaired listeners. Design: A prospective loudness matching experiment utilizing the method of adjustment. Study Sample: 19 consenting native English speakers (Mean age: 61.4, SD: 16.4) with bilateral mild to moderate high-frequency sensorineural hearing loss, indicative of presbycusis. Stimuli: 13 vowel-consonant-vowel (VCV) nonsense syllables, exclusively utilizing the flanking vowel /u/. Results: Descriptive analysis revealed a strong time-order effect influencing loudness judgments for 7 of the 13 VCV test stimuli. Statistical testing showed no significant didference (P = 0.94) between the relative amplitudes corresponding to the point of equal loudness for plosive-containing versus non-plosive-containing VCV stimuli. However, 6 individual VCV stimuli, containing consonants from 4 separate manners of articulation, produced significant loudness matching data (P < 0.01). Conclusions: The results falsify the hypothesis that plosives, analyzed collectively as a class, possess a heavier perceptual loudness weighting than non-plosive consonants. While 6 individual VCV stimuli indicated potential individual consonantal loudness weightings, these findings must be interpreted cautiously due to the restriction to a single vowel context and the presence of procedural time-order biases.
Yildiran, O. F.; Ni, L.; Landy, M. S.
Show abstract
Previous work showed that observers integrate audiovisual duration cues optimally when cue-conflict is small. Does causal inference lead to a breakdown of audiovisual integration when duration conflicts are large? We addressed this by testing a wide range of duration cue-conflicts. Participants compared the auditory durations of a test and a standard stimulus. Audiovisual durations were consistent in the test stimulus, but differed by seven conflict durations (up to 250 ms) in the standard. Two levels of auditory noise were tested. Auditory duration percepts shifted systematically toward the visual duration, especially with high auditory noise. The shift was proportional to cue-conflict magnitude, inconsistent with causal inference. We compared several models. A heuristic model in which the observer probabilistically switches between the visual and auditory cues was preferred for most participants, although performance differences across models were small. Within the tested conflict range, the forced fusion, causal inference, and probabilistic cue switching models produced overlapping, near-linear shifts as a function of cue-conflict. Model simulations further revealed that given the measured sensory noise, forced fusion and causal inference can be discriminated only with unreasonably large conflicts. Together, while our results suggest that observers do not rely on causal inference when judging auditory durations under our conditions, high sensory encoding noise in auditory duration limits the discriminability of competing computational models.